Back

International Journal of Medical Informatics

Elsevier BV

All preprints, ranked by how well they match International Journal of Medical Informatics's content profile, based on 26 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
Detecting Problematic Opioid Use in the Electronic Health Record: Automation of the Addiction Behaviors Checklist in a Chronic Pain Population

Chatham, A. H.; Bradley, E. D.; Schirle, L.; Sanchez-Roige, S.; Samuels, D. C.; Jeffery, A. D.

2023-06-12 health informatics 10.1101/2023.06.08.23290894 medRxiv
Top 0.1%
22.7%
Show abstract

ImportanceIndividuals whose chronic pain is managed with opioids are at high risk of developing an opioid use disorder. Large data sets, such as electronic health records, are required for conducting studies that assist with identification and management of problematic opioid use. ObjectiveDetermine whether regular expressions, a highly interpretable natural language processing technique, could automate a validated clinical tool (Addiction Behaviors Checklist1) to expedite the identification of problematic opioid use in the electronic health record. DesignThis cross-sectional study reports on a retrospective cohort with data analyzed from 2021 through 2023. The approach was evaluated against a blinded, manually reviewed holdout test set of 100 patients. SettingThe study used data from Vanderbilt University Medical Centers Synthetic Derivative, a de-identified version of the electronic health record for research purposes. ParticipantsThis cohort comprised 8,063 individuals with chronic pain. Chronic pain was defined by International Classification of Disease codes occurring on at least two different days.18 We collected demographic, billing code, and free-text notes from patients electronic health records. Main Outcomes and MeasuresThe primary outcome was the evaluation of the automated method in identifying patients demonstrating problematic opioid use and its comparison to opioid use disorder diagnostic codes. We evaluated the methods with F1 scores and areas under the curve - indicators of sensitivity, specificity, and positive and negative predictive value. ResultsThe cohort comprised 8,063 individuals with chronic pain (mean [SD] age at earliest chronic pain diagnosis, 56.2 [16.3] years; 5081 [63.0%] females; 2982 [37.0%] male patients; 76 [1.0%] Asian, 1336 [16.6%] Black, 56 [1.0%] other, 30 [0.4%] unknown race patients, and 6499 [80.6%] White; 135 [1.7%] Hispanic/Latino, 7898 [98.0%] Non-Hispanic/Latino, and 30 [0.4%] unknown ethnicity patients). The automated approach identified individuals with problematic opioid use that were missed by diagnostic codes and outperformed diagnostic codes in F1 scores (0.74 vs. 0.08) and areas under the curve (0.82 vs 0.52). Conclusions and RelevanceThis automated data extraction technique can facilitate earlier identification of people at-risk for, and suffering from, problematic opioid use, and create new opportunities for studying long-term sequelae of opioid pain management. Key PointsO_ST_ABSQuestionC_ST_ABSCan an interpretable natural language processing method automate a valid, reliable clinical tool in order to expedite the identification of problematic opioid use in the electronic health record? FindingsIn this cross-sectional study of patients with chronic pain, an automated natural language processing approach identified individuals with problematic opioid use that were missed by diagnostic codes. MeaningRegular expressions can be used in automatically identifying problematic opioid use in an interpretable and generalizable manner.

2
Can longitudinal electronic health record data identify patients at higher risk of developing long COVID?

Shanmugam, P.; Bair, M.; Pendl-Robinson, E.; Hu, X. C.

2024-02-09 health informatics 10.1101/2024.02.08.24302528 medRxiv
Top 0.1%
22.1%
Show abstract

With hundreds of millions of COVID-19 infections to date, a considerable portion of the population has developed or will develop long COVID. Understanding the prevalence, risk factors, and healthcare costs of long COVID can be of significant societal importance. To investigate the utility of large-scale electronic health record (EHR) data in identifying and predicting long COVID, we analyzed a sample of 1.23 million participants from the National COVID Cohort Collaborative (N3C), a longitudinal EHR data repository from 80 sites in the US with over 8 million COVID-19 patients. We characterized the prevalence of long COVID using a few different types of definitions to illustrate their relative strengths and weaknesses. Then we developed machine learning models to predict the risk of developing long COVID using demographic factors and comorbidity in the EHR. The risk factors for long COVID include patient age; sex; smoking status; and comorbidities characterized by the Charlson Comorbidity Index (CCI). We were able to predict three types of long COVID with low to moderate levels of accuracy (AUC 0.599 - 0.734). We found that age and CCI were most predictive of long COVID diagnosis. Ongoing work includes applying the fair machine learning framework to the long COVID predictive models. We are implementing fairness and bias mitigation methods to model fitting through the following steps, selecting fairness metrics, preparing data and model, evaluating fairness metrics, applying bias mitigation methods to the dataset, and comparing model results and fairness metrics before and after the mitigation. The objective is to achieve equalized odds, a statistical notion that ensures classification algorithms do not discriminate against protected groups (such as sex and race/ethnicity). Results from the fairness-based machine learning will be included in the conference presentation.

3
Development and Temporal Evaluation of Multimodal Machine Learning Models to Predict High Inpatient Opioid Exposure

Kale, S.; Singh, D.; Truumees, E.; Geck, M.; Stokes, J.

2026-04-02 health informatics 10.64898/2026.03.31.26349842 medRxiv
Top 0.1%
18.9%
Show abstract

High inpatient opioid exposure is associated with increased risk of persistent opioid use. Early identification of high-risk patients may improve opioid stewardship. We developed machine learning models to predict high opioid exposure during hospitalization using electronic health record data from MIMIC-IV. We conducted a retrospective study of 223,452 unique first hospital admissions in MIMIC-IV. The outcome was high opioid exposure, defined as the top decile among opioid-exposed admissions (MME/day [≥] 225), representing 2.65% of all admissions. Structured early-admission features included demographics, admission characteristics, laboratory utilization and abnormality summaries, and 24-hour procedural indicators. Discharge-note data were incorporated using ClinicalBERT embeddings and interpretable bigram features. Models were trained using an 80/10/10 split and evaluated with temporal validation on the most recent 10% of admissions. Performance was assessed using ROC-AUC and PR-AUC with 95% confidence intervals. Among structured-only models, XGBoost achieved the best test performance (ROC-AUC 0.932 [0.924-0.940]; PR-AUC 0.223 [0.193-0.262]). The combined structured and notes model improved precision-recall performance (ROC-AUC 0.932 [0.920-0.943]; PR-AUC 0.276 [0.229-0.331]). Temporal evaluation showed similar discrimination (ROC-AUC 0.929; PR-AUC 0.223). High-risk bigrams included procedural terms such as "external fixation" and "cervical discectomy." Integration of structured and text-derived features improved risk stratification compared to structured data alone. Interpretable bigram signals reflected procedural complexity and orthopedic pathology, reinforcing the clinical plausibility of model predictions. Multimodal EHR-based models accurately predict high inpatient opioid exposure and may support targeted opioid stewardship during hospitalization.

4
Algorithmic Identification of Potentially High Risk Abdominal Presentations (PHRAPs) to the Emergency Department: A Clinically-Oriented Machine Learning Approach

Kuzma, R.; Saraswathula, V.; Moon, K.; Kelz, R. R.; Friedman, A. B.

2022-02-09 emergency medicine 10.1101/2022.02.08.22270691 medRxiv
Top 0.1%
18.7%
Show abstract

BackgroundOlder adults presenting to emergency departments (EDs) with abdominal pain have been shown to be at high risk of subsequent morbidity and mortality. Yet, such presentations are poorly studied in national databases. Claims databases do not record the patients symptoms at the time of presentation to the ED, but rather the diagnosis after testing and evaluation, limiting study of care and outcomes for these high risk abdominal presentations. ObjectivesWe sought to develop an algorithm to define a patient population with potentially high risk abdominal presentations (PHRAPs) using only variables commonly available in claims data. Research DesignTrain a machine learning model to predict abdominal pain chief complaints using the National Hospital Ambulatory Medical Care Survey (NHAMCS), a nationally-representative database of abstracted ED medical records. SubjectsAll patients contained in NHAMCS data from 2013-2018. 2013-2017 were used for predictive modeling and 2018 was used as a hold-out test set. MeasuresPositive predictive value and sensitivity of the predictive algorithm against a hold-out test set of NHAMCS patients the algorithm was blinded to during training. Predictions were assessed for agreement with either a chief complaint of abdominal pain (contained in "Reason for Visit 1"), or an expanded definition intended to capture visits which were for abdominal concerns. These included secondary or tertiary complaints of abdominal pain or other abdominal conditions, other abdominal-related chief complaint (e.g. nausea or diarrhea, but not pain), discharge diagnosis of an abdominal condition, or reception of an abdominal CT or ultrasound. ResultsAfter validation on a hold-out data set, a gradient boosting machine (GBM) was the best best-performing machine learning model, but a logistic regression model had similar performance and may be more explainable and useful to future researchers. The GBM predicted a chief complaint of abdominal pain with a positive predictive value of 0.60 (95% CI of 0.56, 0.64) and a sensitivity of 0.29 (95% CI of (0.27, 0.32). Nearly all false positives still exhibited signs of "abdominal concerns" for patients: using the expanded definition of "abdominal concern" the model had a PPV of >0.99 (95% CI of 0.99, 1.00) and sensitivity of 0.12 (95% CI of 0.11, 0.13). ConclusionThe algorithm we report defines a patient population with abdominal concerns for further study of treatment and outcomes to inform the development of clinical pathways.

5
Identify Patients at Risk of HIV Using a Clinical Large Language Model from Electronic Health Records

Liu, Y.; Chen, Z.; Suman, P.; Cho, H.; Prosperi, M.; Wu, Y.

2026-04-23 hiv aids 10.64898/2026.04.21.26351427 medRxiv
Top 0.1%
18.7%
Show abstract

This study developed a large language model (LLM)-based solution to identify people at HIV risk using electronic health records. We transformed structured EHR data, including demographics, diagnoses, and medications, into narrative descriptions ordered by visit date and applied GatorTron, a widely used clinical LLM trained on 82 billion words of de-identified clinical text. We compared GatorTron with traditional machine learning models, including LASSO and XGBoost. We identified a cohort with 54,265 individuals, where only 3,342 (6%) had new HIV diagnoses. Our LLM solution, based on GatorTron, achieved excellent performance, reaching an F1 score of 53.5% and an AUC of 0.88, comparable to traditional machine learning approaches. Subgroup analysis showed that, across age, sex, and race/ethnicity groups, both LLM and traditional models achieved AUCs above 0.82. Interpretability analyses showed broadly consistent patterns across LLM models and traditional machine learning models.

6
Building Prediction Models for 30-Day Readmissions Among ICU Patients Using Both Structured and Unstructured Data in Electronic Health Records

Moerschbacher, A.; He, Z.

2021-08-11 health informatics 10.1101/2021.08.10.21261858 medRxiv
Top 0.1%
18.4%
Show abstract

ICU readmissions are associated with poor outcomes for patients and poor performance of hospitals. Patients who are readmitted have an increased risk of in-hospital deaths; hospitals with a higher readmission rate have a reduced profitability, due to an increase in cost and reduced payments from Medicare and Medicaid programs. Predicting a patients likelihood of being readmitted to the ICU can help reduce early discharges, the risk of in-hospital deaths, and help increase profitability. In this study, we built and evaluated multiple machine learning models to predict 30-day readmission rates of ICU patients in the MIMIC-III database. We used both the structured data including demographics, laboratory tests, comorbidities, and unstructured discharge summaries as the predictors and evaluated different combinations of features. The best performing model in this study Logistic Regression achieved an AUROC of 75.7%. This study shows the potential of leveraging machine learning and deep learning for predicting ICU readmissions.

7
Computational Phenotypes for Patients with Opioid-Related Disorders Presenting to the Emergency Department

Gilson, A.; Schulz, W. L.; Lopez, K.; Young, P.; Pandya, S.; Coppi, A.; Chartash, D.; Fiellin, D.; D'Onofrio, G.; Taylor, R. A.

2023-03-29 emergency medicine 10.1101/2023.03.24.23287638 medRxiv
Top 0.1%
15.5%
Show abstract

ObjectiveWe aimed to discover computationally-derived phenotypes of opioid-related patient presentations to the emergency department (ED) via clinical notes and structured electronic health record (EHR) data. MethodsThis was a retrospective study of ED visits from 2013-2020 across ten sites within a regional healthcare network. We derived phenotypes from visits for patients 18 years of age with at least one prior or current documentation of an opioid-related diagnosis. Natural language processing was used to extract clinical entities from notes, which were combined with structured data within the EHR to create a set of features. We performed Latent Dirichlet allocation to identify topics within these features. Groups of patient presentations with similar attributes were identified by cluster analysis. ResultsIn total 82,577 ED visits met inclusion criteria. The 30 topics discovered ranged from those related to substance use disorder, chronic conditions, mental health, and medical management. Clustering on these topics identified nine unique cohorts with one-year survivals ranging from 84.2-96.8%, rates of one-year ED returns from 9-34%, rates of one-year opioid event 10-17%, rates of medications for opioid use disorder from 17-43%, and a median Carlson comorbidity index of 2-8. Two cohorts of phenotypes were identified related to chronic substance use disorder, or acute overdose. ConclusionsOur results indicate distinct phenotypic clusters with varying patient-oriented outcomes which provide future targets better allocation of resources and therapeutics. This highlights the heterogeneity of the overall population, and the need to develop targeted interventions for each population.

8
Machine Learning Methods to Predict Survival in Patients Following Traumatic Aortic Injury

Shiban, N.; Gaul, J.; Zhan, H.; Elhabr, A.; Kokabi, N.; Johnson, J.-O.; Hanna, T.; Schrager, J.; Gichoya, J.; Banerjee, I.; Trivedi, H.

2021-07-31 health informatics 10.1101/2021.07.28.21261166 medRxiv
Top 0.1%
15.3%
Show abstract

The National Trauma Data Bank (NTDB) is a resource of diagnostic, treatment, and outcomes information in trauma patients. We leverage the NTDB and machine learning techniques to predict survival following traumatic aortic injury. We create two predictive models using the NTDB - 1) using all data and, 2) using only data available in the first hour after arrival (prospective data). Seven discriminative model types were tested before and after feature engineering to reduce dimensionality. The top performing model was XGBoost, achieving an overall accuracy of 0.893 using all data and 0.855 using prospective data. Feature engineering improved performance of all models. Glasgow Coma Scale score was the most important factor for survival, and thoracic endovascular aortic repair was more common in patients that survived. Smoking, pneumonia, and urinary tract infection predicted poor survival. We also note concerning disparities in outcomes for black and uninsured patients that may reflect differences in care.

9
Development and Evaluation of Machine Learning Models for the Detection of Emergency Department Patients with Opioid Misuse from Clinical Notes

Shahid, U.; Parde, N.; Smith, D. L.; Dickinson, G.; Bianco, J.; Thorpe, D.; Hota, M.; Afshar, M.; Karnik, N. S.; chhabra, n.

2024-12-12 emergency medicine 10.1101/2024.12.11.24318875 medRxiv
Top 0.1%
15.2%
Show abstract

ObjectivesThe accurate identification of Emergency Department (ED) encounters involving opioid misuse is critical for health services, research, and surveillance. We sought to develop natural language processing (NLP)-based models for the detection of ED encounters involving opioid misuse. MethodsA sample of ED encounters enriched for opioid misuse was manually annotated and clinical notes extracted. We evaluated classic machine learning (ML) methods, fine-tuning of publicly available pretrained language models, and a previously developed convolutional neural network opioid classifier for use on hospitalized patients (SMART-AI). Performance was compared to ICD-10-CM codes. Both raw text and text transformed to the United Medical Language System were evaluated. Face validity was evaluated by term feature importance. ResultsThere were 1123 encounters used for training, validation, and testing. Of the classic ML methods, XGBoost had the highest AU_PRC (0.936), accuracy (0.887), and F1 score (0.863) which outperformed ICD-10-CM codes [accuracy 0.870; F1 0.830]. Logistic regression, support vector machine, and XGBoost models had higher AU_PRC using transformed text, while decision trees performed better using raw text. Excluding XGBoost, fine-tuned pre-trained language models outperformed classic ML methods. The best performing model was the fine-tuned SMART-AI based model with domain adaptation [AU_PRC 0.948; accuracy 0.882; F1 0.851]. Explainability analyses showed the most predictive terms were heroin, opioids, alcoholic intoxication, chronic, cocaine, opiates, and suboxone. ConclusionsNLP-based models outperform entry of ICD-10-CM diagnosis codes for the detection of ED encounters with opioid misuse. Fine tuning with domain adaptation for pre-trained language models resulted in improved performance.

10
Differentiation of Fungal, Viral, and Bacterial Sepsis using Multimodal Deep Learning

Boussina, A.; Ramesh, K.; Arora, H.; Ratadiya, P.; Nemati, S.

2023-04-11 health informatics 10.1101/2023.04.10.23288378 medRxiv
Top 0.1%
15.2%
Show abstract

Sepsis is a major cause of morbidity and mortality worldwide, and is caused by bacterial infection in a majority of cases. However, fungal sepsis often carries a higher mortality rate both due to its prevalence in immunocompromised patients as well as delayed recognition. Using chest x-rays, associated radiology reports, and structured patient data from the MIMIC-IV clinical dataset, the authors present a machine learning methodology to differentiate between bacterial, fungal, and viral sepsis. Model performance shows AUCs of 0.81, 0.83, 0.79 for detecting bacterial, fungal, and viral sepsis respectively, with best performance achieved using embeddings from image reports and structured clinical data. By improving early detection of an often missed causative septic agent, predictive models could facilitate earlier treatment of non-bacterial sepsis with resultant associated mortality reduction.

11
Stigmatizing Language Detection in Opioid Use Disorder Patient-Directed Discharge Clinical Documentation: A Privacy-Preserving Analysis Using a Locally Deployed Large Language Model

Izzo, J. A.; McIntyre, A. M.; Nguyen, J.; Bashaw, D.; Torrance, C. A.; Foster, J.

2026-06-01 health informatics 10.64898/2026.05.29.26354402 medRxiv
Top 0.1%
15.1%
Show abstract

Objective: Stigmatizing language in the electronic health record (EHR) has been associated with adverse patient experience in substance use disorder care, including opioid use disorder (OUD). This study evaluated a privacy-preserving, locally-deployed large language model as a method to detect stigmatizing language documentation in OUD patients with patient-directed discharge (PDD). Methods: A retrospective cohort study of 477 inpatient admissions from the MIMIC-IV database with a diagnosis of opioid use disorder were classified using a locally deployed Gemma-4-31b-it-bf16 model and predefined 140 term lexicon to identify stigmatizing language in clinical documentation. Results: Analysis of clinical documentation showed stigmatizing language was present in 84.1% (190/226) in the PDD cohort vs 62.2% (156/251) in the non-PDD cohort, with an unadjusted odds ratio of 3.21 (95% CI 2.07-4.98; p < 0.0001). After adjustment for age, sex, insurance status, marital status, and race, PDD discharge remained an independent predictor of stigmatizing documentation (aOR 2.24, 95% CI 1.40-3.59; p < 0.0001). Further analysis of stigma intensity showed higher stigmatizing markers in the PDD cohort vs the non-PDD cohort (2.85 {+/-} 2.39 vs 2.02 {+/-} 2.44; p < 0.0001). Discussion and Conclusion: Stigmatizing language is detected with increased frequency and prevalence in clinical documentation of OUD patients that initiate PDD compared to those that adhere to standard discharge processes. A locally deployed large language model (LLM) offers a scalable, privacy-preserving method to audit clinical documentation for stigmatizing language.

12
Predicting 30 Days Hospital Readmission for Heart Failure patients using word embeddings

Shakya, P. R.; Khaneja, A.; Wagholikar, K. B.

2025-02-10 health informatics 10.1101/2025.02.07.25321871 medRxiv
Top 0.1%
15.0%
Show abstract

Heart Failure (HF) is a public health concern with a wider impact on quality of life and cost of care. One of the major challenges in HF is the higher rate of unplanned readmissions and sub-optimal performance of models to predict the readmissions. Hence, in this study, we implemented embeddings-based approaches to generate features for improving model performance. Specifically, we compared three embedding approaches including word2vec on terminology codes and CUIs, and BERT on concept descriptions with baseline (one hot-encoding). We found that the embedding approaches significantly improved the performance of the prediction models, and word2vec on the study dataset outperformed pre-trained BERT model.

13
Development of a natural language processing application to extract and categorize mentions of violence from mental healthcare records text

Li, L.; Sondh, S.; Sondh, H. K.; Stewart, R.; Roberts, A.

2026-03-26 health informatics 10.64898/2026.03.22.26348435 medRxiv
Top 0.1%
14.9%
Show abstract

BackgroundExperiences of violence are reported frequently by mental health service users, victims of violence are at a greater risk of mental health disorders, and violence may sometimes occur as a consequence of a mental disorder. Electronic health records (EHRs) are an important source of information about healthcare, and its social context. Occurrences of violence are not routinely recorded as structured data in EHRs but are however recorded in the free text narrative. ObjectiveOur objective was to address this research gap by creating a natural language processing (NLP) application that extracts information related to various forms of violence (physical (non-sexual), sexual, emotional, and financial) from the EHR of a large south London mental health service. Additionally, we aimed to extract features concerning the patients role (victimization vs. perpetration), timing (recent vs. historic), domestic context, presence (actual, threat, or unclear), and polarity (affirmed, abstract, or negated) of the violent behaviors. MethodsTwo raters independently annotated 6,500 randomly selected segments of clinical notes containing violence-related keywords from a large mental healthcare provider in South London, each containing 400 characters (with approximately 200 characters before and after the keyword) after rigorous training using a pre-defined and approved coding book provided by senior professionals. We utilized 90% of the annotated data for fine-tuning a multi-label BERT model (employing 5-fold cross-validation) with the remaining 10% of data reserved for a blind test. ResultsThe model performed well on the blind test set for emotional violence (F1= 0.89), financial violence (0.88), physical (non-sexual) violence (0.84), and unspecified violence (0.81), and the patient role (0.89 as perpetrator; 0.84 as victim), polarity (0.89 for affirmed behavior), presence (0.95 for actual violence), and domestic settings (0.88). We were unable to achieve satisfactory results in capturing temporal aspects (0.65 for past violence). ConclusionsWe were able to improve substantially on previously developed NLP for ascertaining violence in routine mental health records, providing novel opportunities for both surveillance and research.

14
Toward a National Registry for Inborn Errors of Immunity in Peru: A Qualitative Implementation Study

Veramendi-Espinoza, L. E.; De la Cruz-Torralva, K.; Pezo-Pezo, A.; Vargas-Herrera, J. R.; Neyra Quijandria, J.; Martina, M.; Sullivan, K. E.; Huaman, M. A.; Knapke, J. M.

2026-06-15 allergy and immunology 10.64898/2026.06.12.26355539 medRxiv
Top 0.1%
13.7%
Show abstract

Background: Peru lacks an integrated information system for patients with Inborn Errors of Immunity (IEI). Although disease registries are essential tools for data management and health planning, their success depends on implementation science approaches that account for local contextual factors. This study reports Phase I of a three-phase mixed-methods implementation project to design and develop a national IEI registry. Methods: Phase I consisted of a phenomenological qualitative study exploring stakeholder perspectives. Semi-structured focus groups and in-depth interviews were conducted with 29 key stakeholders across four groups: policy-makers, clinical experts, end-users (immunologists, residents, allied health personnel), and patient organization representatives. Interviews followed a guide structured around four a priori domains (structure, navigation, feasibility, and perception of existing systems). Discussions were conducted in Spanish, audio-recorded, transcribed verbatim, and coded using ATLAS.ti. A hybrid thematic analysis combining deductive and inductive coding was performed. Data elements proposed for the registry were triangulated with qualitative findings. Results: Thirty-six initial codes were consolidated into 15 categories, which were further integrated into four overarching themes conceptualized as pathways toward intention to use: (1) Environment, where governance, regulatory backing, and sustainable financing were identified as key enablers, while limited interoperability emerged as a structural barrier; (2) Technical Dimension, emphasizing usability, alignment with clinical workflow, and a hierarchical data architecture (demographic, clinical, therapeutic); (3) Users, highlighting clinical leadership, protected time, digital readiness, and perceived usefulness as stronger motivators than financial incentives; and (4) Patients, underscoring data protection, transparency, trust, and advocacy as essential for legitimacy and sustainability. Conclusions: A national IEI registry in Peru is perceived as necessary and feasible if implemented with strong regulatory foundations, interoperable design, robust data security, and user-centered architecture. These findings informed the development of an initial functional prototype and the operational plan for Phase II, focused on usability evaluation.

15
Development of a federated learning approach to predict acute kidney injury in adult hospitalized patients with COVID-19 in New York City

Jaladanki, S. K.; Vaid, A.; Sawant, A. S.; Xu, J.; Shah, K.; Dellepiane, S.; Paranjpe, I.; Chan, L.; Kovatch, P.; Charney, A.; Wang, F.; Glicksberg, B. S.; Singh, K.; Nadkarni, G. N.

2021-07-28 nephrology 10.1101/2021.07.25.21261105 medRxiv
Top 0.1%
13.2%
Show abstract

Federated learning is a technique for training predictive models without sharing patient-level data, thus maintaining data security while allowing inter-institutional collaboration. We used federated learning to predict acute kidney injury within three and seven days of admission, using demographics, comorbidities, vital signs, and laboratory values, in 4029 adults hospitalized with COVID-19 at five sociodemographically diverse New York City hospitals, between March-October 2020. Prediction performance of federated models was generally higher than single-hospital models and was comparable to pooled-data models. In the first use-case in kidney disease, federated learning improved prediction of a common complication of COVID-19, while preserving data privacy.

16
Representing Injuries in Trauma Patients: Development and Evaluation of Embeddings for Injuries

Szolnoky, K.; Attergrim, J.; Ashfaq, A.; Linusson, H.; Gerdin Wärnberg, M.; Berg, J.

2026-01-06 emergency medicine 10.64898/2026.01.03.26343379 medRxiv
Top 0.1%
13.1%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWO_ST_ABSBackgroundC_ST_ABSTrauma patients present with heterogeneous injury patterns that are challenging to represent in statistical models. Traditional approaches either use high-dimensional one-hot encoding, resulting in sparse features, or aggregate injuries into summary scores that lose patient-specific detail. This study developed data-driven ICD-10 embeddings for trauma injuries and evaluated their ability to preserve injury information. MethodsUsing the National Trauma Data Bank, we trained autoencoder models on all trauma patients from 2018 to generate dense vector representations of ICD-10 injury codes. We evaluated embeddings of dimensions 2, 4, 8, 16, and 32 against one-hot encoding using three prediction tasks: in-hospital mortality, emergency department disposition, and blood transfusion within 24 hours. For each hospital included, we trained separate logistic regression and LightGBM models using 2018 data from that hospital, then evaluated performance on 2019 data from the same hospital. Performance was measured using area under the receiver operating characteristic curve (AUC) and stratified by hospital size. ResultsIn LightGBM models, 8-dimensional embeddings improved AUC compared to one-hot encoding of 0.08 (95% CI: 0.06, 0.10) in small hospitals, 0.03 (0.02, 0.04) in medium hospitals, and 0.02 (0.01, 0.02) in large hospitals, with comparable performance in major hospitals (0.00 [-0.01, 0.01]). In logistic regression, 32-dimensional embeddings showed AUC improvements of 0.03 (0.01, 0.05), 0.02 (0.01, 0.03), and 0.02 (0.02, 0.03) for small, medium, and large hospitals respectively, with similar performance in major hospitals (0.01 [0.00, 0.01]). ConclusionICD-10 code injury embeddings with [&ge;]8 dimensions preserve clinically relevant information and can outperform one-hot encoding while reducing dimensionality. The embeddings and software are openly available to support further trauma research and applications.

17
An application of machine learning to assist medication order review by pharmacists in a health care center

Thibault, M.; Lebel, D.

2019-11-27 health informatics 10.1101/19013029 medRxiv
Top 0.1%
13.0%
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWThe objective of this study was to determine if it is feasible to use machine learning to evaluate how a medication order is contextually appropriate for a patient, in order to assist order review by pharmacists. A neural network was constructed using as input the sequence of word2vec embeddings of the 30 previous orders, as well as the currently active medications, pharmacological classes and ordering department, to predict the next order. The model was trained with data from 2013 to 2017, optimized using 5-fold cross-validation, and tested on orders from 2018. A survey was developed to obtain pharmacist ratings on a sample of 20 orders, which were compared with predictions. The training set included 1 022 272 orders. The test set included 95 310 orders. Baseline training set top 1, top 10 and top 30 accuracy using a dummy classifier were respectively 4.5%, 23.6% and 44.1%. Final test set accuracies were, respectively, 44.4%, 69.9% and 80.4%. Populations in which the model performed the best were obstetrics and gynecology patients and newborn babies (either in or out of neonatal intensive care). Pharmacists agreed poorly on their ratings of sampled orders with a Fleiss kappa of 0.283. The breakdown of metrics by population showed better performance in patients following less variable order patterns, indicating potential usefulness in triaging routine orders to less extensive pharmacist review. We conclude that machine learning has potential for helping pharmacists review medication orders. Future studies should aim at evaluating the clinical benefits of using such a model in practice.

18
Leveraging Machine Learning for Developing and Validating a Neonatal Acute Kidney Injury Prediction Model (NEPHRO): A Comprehensive Evidence-Based Neonatal AKI Risk Stratification Tool

Mohamed, T. H.; Bambach, S.; Spencer, J. D.; Rust, L.; Patel, S.; magers, j.; Neyra, J.; Wilson, F. P.; Ning, X.; Newland, J.; Rust, S.; Slaughter, J. L.

2025-06-13 nephrology 10.1101/2025.06.12.25329508 medRxiv
Top 0.1%
13.0%
Show abstract

BackgroundAcute kidney injury (AKI) is a serious and common complication among critically ill neonates. Preventing or treating AKI early requires timely prediction, but current tools to forecast AKI in neonates are limited. We proposed that machine learning could help predict AKI by analyzing routinely collected clinical data. MethodsWe conducted a retrospective analysis of 8,059 critically ill neonates admitted to our level IV neonatal intensive care unit (NICU) from January 2017 to December 2021 for development and cross-validation (n=5,443, 68%), and from January 2022 to June 2024 (n=2,616, 32%) for temporally validating a neonatal AKI predictive model. Risk factors for model input were identified from the literature, and data were extracted from electronic health records. According to the neonatal modification of kidney disease: Improving Global Outcomes criteria, an increase in serum creatinine (SCr) and/or a decrease in urine output (UOP) defines neonatal AKI. We trained a machine learning model using a least absolute shrinkage and selection operator (LASSO) algorithm to develop and validate a predictive AKI model. The area under the receiver operating characteristic curve (AUROC) and F1-scores evaluated the models performance. FindingsAmong 206,220 NICU patient days, AKI occurred in 881 (11%) neonates. Using 27 potential AKI risk factors, LASSO identified key AKI predictors: fluid balance, hypotension requiring vasopressors, invasive ventilation, sepsis, surgical procedures, and congenital kidney and urinary tract anomalies. The model predicted the occurrence of a critical SCr increase or UOP decrease over the next 48 hours, with an AUROC of 0.814 (95% CI: 0.787-0.843) in development and 0.815 (0.795-0.834) in validation datasets. InterpretationWe developed a machine learning-based model that reliably predicted neonatal AKI before it became clinically apparent by conventional parameters. By identifying high-risk neonates earlier, timely interventions can be deployed to improve outcomes.

19
Developing and optimizing machine learning algorithms for predicting in-hospital patient charges for Congestive Heart Failure Exacerbations, Chronic Obstructive Pulmonary Disease Exacerbations and Diabetic Ketoacidosis

Arnold, M. C.; Boland, M. R.; Liou, L.

2023-12-18 health informatics 10.1101/2023.12.17.23298944 medRxiv
Top 0.1%
12.9%
Show abstract

BackgroundHospitalizations for exacerbations of congestive heart failure (CHF), chronic obstructive pulmonary disease (COPD) and diabetic ketoacidosis (DKA) are costly in the United States. ObjectiveThe purpose of this study is to predict in-hospital charges for each condition using Machine Learning (ML) models. MethodsWe conducted a retrospective cohort study on national discharge records of hospitalized adult patients from January 1st, 2016, to December 31st, 2019. We used numerous ML techniques to predict in-hospital total cost. ResultsWe found that linear regression (LM), gradient boosting (GBM) and extreme gradient boosting (XGB) models had good predictive performance and were statistically equivalent, with training R-Squared values ranging from 0.49-0.95 for CHF; 0.56-0.95 for COPD; and 0.32-0.99 for DKA. We identified important key features driving costs, including patient age, length-of-stay, number of procedures. and elective/non-elective admission. ConclusionsML methods may be used to accurately predict costs and identify drivers of high cost for COPD exacerbations, CHF exacerbations and DKA. Overall, our findings may inform future studies that seek to decrease the underlying high patient costs for these conditions.

20
An Explainable Machine Learning Framework for Predicting the Risk of Buprenorphine Treatment Discontinuation for Opioid Use Disorder

Faysal, J. A.; E Alam, M. N.; Young, G. J.; Lo-Ciganic, W.-H.; Goodin, A. J.; Huang, J. L.; Wilson, D. L.; Park, T. W.; Hasan, M. M.

2023-11-03 health informatics 10.1101/2023.11.02.23297982 medRxiv
Top 0.1%
12.8%
Show abstract

ObjectivesBuprenorphine is an effective evidence-based medication for opioid use disorder (OUD). Yet premature discontinuation undermines treatment effectiveness, increasing risk of mortality and overdose. We developed and evaluated a machine learning (ML) framework for predicting buprenorphine care discontinuity within 1-year following treatment initiation. MethodsThis retrospective study used United States 2018-2021 MarketScan commercial claims data of insured individuals aged 18-64 who initiated buprenorphine between July 2018 and December 2020 with no buprenorphine prescriptions in the previous six months. We measured buprenorphine prescription discontinuation gaps of [&ge;]30 days within the first year of initiating treatment. We developed predictive models employing logistic regression, decision tree classifier, random forest, XGBoost, Adaboost, and random forest-XGBoost ensemble. We applied recursive feature elimination with cross-validation to reduce dimensionality and identify the most predictive features while maintaining model robustness. We focused on two distinct treatment stages: at the time of treatment initiation and one and three months after treatment initiation. We employed SHapley Additive exPlanations (SHAP) analysis that helped us explain the contributions of different features in predicting buprenorphine discontinuation. We stratified patients into risk subgroups based on their predicted likelihood of treatment discontinuation, dividing them into decile subgroups. Additionally, we used a calibration plot to analyze the reliability of the models. ResultsA total of 30,373 patients initiated buprenorphine and 14.98% (4,551) discontinued treatment. C-statistic varied between 0.56 and 0.76 for the first-stage models including patient-level demographic and clinical variables. Inclusion of proportion of days covered (PDC) measured at one-month and three-month following treatment initiation significantly increased the models discriminative power (C-statistics: 0.60 to 0.82). Random forest (C-statistics: 0.76, 0.79 and 0.82 with baseline predictors, one-month PDC and three-month PDC, respectively) outperformed other ML models in discriminative performance in all stages (C-statistics: 0.56 to 0.77). Most influential risk factors of discontinuation included early stage medication adherence, age, and initial days of supply. ConclusionML algorithms demonstrated a good discriminative power in identifying patients at higher risk of buprenorphine care discontinuity. The proposed framework may help healthcare providers optimize treatment strategies and deliver targeted interventions to improve buprenorphine care continuity.